Skip to content

fix(llm_judge): use last-match for verdict extraction in single and pairwise judging - #3921

Open
AUTHENSOR wants to merge 1 commit into
lm-sys:mainfrom
AUTHENSOR:fix/judge-last-match-verdict
Open

AUTHENSOR wants to merge 1 commit into
lm-sys:mainfrom
AUTHENSOR:fix/judge-last-match-verdict

Conversation

@AUTHENSOR

Copy link
Copy Markdown

Summary

The LLM judge in fastchat/llm_judge/common.py extracts verdicts using first-match parsing:

  1. run_judge_single (line 176): re.search(one_score_pattern, judgment) returns the first [[N]] match. If the judge reasons about scores before its final verdict, the first match may be from reasoning text, not the verdict.

  2. run_judge_pair (line 283): if "[[A]]" in judgment substring checks in A-before-B-before-C order. A response containing [[A]] in reasoning and [[B]] as the verdict selects A (first match).

Fix

  • Single: Replace re.search with re.findall and take the last match (matches[-1]).
  • Pairwise: Replace the in substring chain with re.findall(r"\[\[([ABC])\]\]", judgment.upper())[-1].
  • Two-score ([[rating_a,rating_b]]): Same findall + last-match treatment.

This matches the convention used by other eval frameworks (inspect_ai DEFAULT_GRADE_PATTERN uses a greedy prefix to bind to the last verdict; openai/evals cot_classify reverses lines).

Verification

  • black --check: clean (no formatting changes needed)
  • Diff touches only the three extraction blocks (16 lines changed)

…ngle/pair

run_judge_single used re.search (first-match) for [[rating]] extraction.
run_judge_pair used substring 'in' checks (A checked before B before C).
Both allow an early verdict token in reasoning text to override the
judge's final verdict. Switch to findall + last-match.
@shaurya416

Copy link
Copy Markdown

I checked this against run_judge_pair's [[A]]-format branch: a [[A]] token appearing in the judge's reasoning text can override its actual final verdict, since the current code checks substrings in a fixed A-then-B-then-C order rather than reading the last tag. I extracted just that branch verbatim from main (587d5cf) and from this PR's head (e6579c1) and ran it standalone, no repository code imported:

case main this PR
plain [[A]] / [[B]] / [[C]] (controls) A / B / tie A / B / tie
[[A]] in reasoning, [[B]] is the final verdict A (wrong) B
[[A]] and [[B]] both mentioned in reasoning, final verdict [[C]] A (wrong) tie
no bracketed tag at all error error
lowercase [[a]] error (silently missed, since the substring check is case-sensitive) A (the .upper() call in this PR picks it up too)

So the reported defect is fixed, the two controls are unchanged, and the case-insensitive matching this PR adds as a side effect (from judgment.upper()) also recovers a lowercase-tag reply that main currently drops to error. I did not run the single-score or two-score branches this PR also touches, or the repository's own tests.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants